<div> RuLegalNER: a new dataset for Russian legal named entities recognition</div> Open database of scientific publications ITMO UNIVERSITY

RuLegalNER: a new dataset for Russian legal named entities recognition

Journal

Scientific and Technical Journal of Information Technologies, Mechanics and Optics

Shaheen Zein, Mouromtsev Dmitry I., Postny Ignat

UDK004.912

Issue:4 (150)

Download PDF0 Kbyte

Annotation

We address the scarcity of datasets specifically tailored for legal NER in the Russian language and investigate the generalization capabilities of models towards unseen named entities. A rule-based program developed by legal experts at Tag-Consulting Company was employed to automatically annotate legal texts and create the RuLegalNER dataset. Part of the named entities only exists in the development and test splits, and they are unseen in the training set. RuBERT was utilized as the base architecture for experimental evaluation. Two different architectural extensions were explored: RuBERT with CRF and RuBERT with adapters. These architectures were used to train and evaluate NER models on the RuLegalNER dataset. Utilize RuLegalNER to train and evaluate legal NER models, enhancing performance in the legal domain and studying generalization on unseen entities. A published version of RuLegalNER is presented with detailed statistics and demonstration of the usefulness of RuLegalNER by evaluating modern architectures.

RuLegalNER: a new dataset for Russian legal named entities recognition

Scientific and Technical Journal of Information Technologies, Mechanics and Optics

Annotation

Keywords

Постоянный URL

Articles in current issue

RuLegalNER: a new dataset for Russian legal named entities recognition

Scientific and Technical Journal of Information Technologies, Mechanics and Optics

Annotation

Keywords

Постоянный URL

Поделиться

Articles in current issue